Repository navigation
Reduce dashboard CPU under sustained metric ingestion - #20736
Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
🚀 Dogfood this PR with:
curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20736Or
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20736" |
Tests selectorSelects the full PR test matrix + all PR-gated jobs (ALL) — run-all fallback: 'benchmarks/Aspire.Dashboard.Benchmarks/TelemetryRepositoryMetricsBenchmarks.cs' is neither Layer-1-owned nor matched by a Layer 2 rule Selection computed for commit |
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The optimized retention logic preserves existing semantics and has focused coverage for capacity, persistence, repeated values, and transaction failure recovery.
Review effort: Balanced
Findings: None
What changed in this PR
Optimizes SQLite metric retention to prevent rising dashboard CPU during sustained ingestion.
Changes:
- Caches per-dimension point counts and performs indexed surplus deletion.
- Adds retention, restart, rollback, and multi-dimension tests.
- Adds and documents a capacity-ingestion benchmark.
| File | Description |
|---|---|
src/Aspire.Dashboard/Otlp/Storage/SqliteTelemetryRepository.Metrics.Writes.cs |
Optimizes metric retention. |
tests/Aspire.Dashboard.Tests/TelemetryRepositoryTests/MetricsTests.cs |
Covers dimension-specific retention. |
tests/Aspire.Dashboard.Tests/TelemetryRepositoryTests/SqliteTelemetryPersistenceTests.cs |
Covers reopening and rollback recovery. |
benchmarks/Aspire.Dashboard.Benchmarks/TelemetryRepositoryMetricsBenchmarks.cs |
Adds capacity-ingestion benchmarking. |
docs/specs/dashboard-persistence.md |
Documents the benchmark. |
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
|
/backport to release/13.6 |
|
Started backporting to |
|
✅ No documentation update needed. Step 5 branch: Triggered signals (2):
Diff confirms no public surface change: the only Retention limits, defaults, and dashboard behavior are already correctly documented on |
…20747) Backport of #20736 to release/13.6 /cc @JamesNK ## Customer Impact In Aspire 13.6, long-running AppHosts with sustained metric ingestion can consume steadily increasing dashboard CPU even under a fixed or idle workload. A reported 15-project AppHost reached approximately 70% of one core after 6.8 hours, with significant SQLite churn and database growth. ## Testing Passed 59 targeted SQLite metrics and persistence tests, excluding quarantined and outerloop tests. The source PR also validated the before/after ingestion benchmarks and smoke-tested all 12 existing metrics-query benchmark cases. Backport PR CI is currently in progress. ## Risk Low. The change is localized to SQLite metric-retention cleanup, preserves the existing retention limit and cascading deletes, and adds focused coverage for retention, database reopening, and transaction-failure recovery. It makes no public API or database schema changes. ## Regression? Yes — introduced in 13.6 with SQLite-backed telemetry in #18924. Co-authored-by: James Newton-King <james@newtonking.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Description
Long-running AppHosts could consume steadily increasing dashboard CPU while ingesting a fixed metrics workload. Metric retention ranked every retained point in each changed dimension on every insertion, including when the dimension was still below its retention limit. This reproduces independently of exemplars.
Cache each dimension's retained point count and load it from SQLite when an existing dimension is first used. Skip retention cleanup below capacity; when capacity is exceeded, use the existing dimension-order index to delete only the oldest surplus points. Existing cascading foreign keys continue to delete the removed points' exemplars and filtered attributes in the same transaction.
Also:
AddHistogramMetricsAtCapacity, preloading 50,000 points per dimension and replaying four exemplars per point. History population and payload construction are outside the measured operation.User-facing usage
Continue starting the AppHost normally, for example with
aspire start. No configuration changes are required. The metric retention limit, one-second development export interval, and exemplar behavior are unchanged; steady-state ingestion no longer scans the retained history on every insertion.Benchmark comparison
Both variants used BenchmarkDotNet 0.15.8, .NET 11 Release builds, Windows x64, the same populated-history workload, three warmup iterations, and ten measured iterations. The baseline used the original retention source from commit
0a42e2c6ae; the fixed variant used this branch's implementation.The baseline used 20 operations per iteration. The fixed path used 512 operations per iteration to avoid short samples. Results are normalized per export, which adds one point per dimension. The fixed five-dimension result had a standard deviation of 1.073 ms, so the speedup figures should be treated as approximate.
Validation
<scratch>below replaces the local temporary directory.git diff --check.Fixes #20725
Checklist
<remarks />and<code />elements on your triple slash comments?